Perplexity2026-10-08 11:58:04Perplexity open-sources retrieval models that let a 0.6B model query indexes built by a 9B modelPerplexity has open-sourced two multimodal embedding models under its pplx-embed-v2-late line, with parameter sizes of 0.6B and 9B. The key design point is vector compatibility between the two: developers can build a database with the 9B model and then run day-to-day searches with the smaller 0.6B model, avoiding the need to use the larger model for every query. The models support retrieval across text, images, and PDF pages. Perplexity said the system keeps a 128-dimensional vector for each token instead of compressing an entire passage into a single vector, a setup aimed at preserving more detail during matching. For PDFs, slide decks, and scanned files, page images can be converted directly into vectors without first extracting text through OCR, which also keeps charts, tables, and layout information. According to the company, both models were trained from the same 18B teacher model and share one vector space. In ViDoRe v3 image retrieval testing, Perplexity reported a score of 62.3% when the 0.6B model handled both indexing and querying. Using the 9B model for indexing and the 0.6B model for querying raised the score to 63.5%, while using the 9B model for both reached 65.2%. The model weights are available on Hugging Face under the MIT license.10